Add opt-in persistent conversation checkpoints - #21
Open
owenqwenstarsky wants to merge 1 commit into
Open
owenqwenstarsky wants to merge 1 commit into
owenqwenstarsky wants to merge 1 commit into
Conversation
Author
|
This PR is not ready to be merged. It remains a draft pending further review and validation, including real-weight 35B testing and additional cache performance measurements. |
owenqwenstarsky
marked this pull request as ready for review
September 12, 2026 15:41
Author
|
If you think this is a good thing to implement I will continue working on it a bit more and testing it out more to get it ready for any real work. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Repeated coding conversations currently recompute their full prompt after each request or process restart. This adds opt-in persistent inference checkpoints so the server and Python engine can restore the longest saved token prefix and process only the unmatched suffix.
The implementation uses a radix index backed by SQLite, immutable shared safetensors KV blocks, and checkpoint-specific recurrent, logits, and prerouter state. It includes namespace fingerprints, checksums, atomic publication, cross-process locking, bounded LRU eviction, exact-message tokenization caching, CLI inspection/clearing, and cached-token usage metrics. Caching defaults to disabled; when enabled, the defaults are 20 GiB and a 2,048-token interval.
Validation:
Publication increased peak MLX allocation by about 24% in this sample. Shared blocks are currently reserialized synchronously, so retained-storage deduplication does not eliminate write work. Real-weight 35B validation, short-prefix break-even measurements, and sustained/cold-filesystem trials remain outstanding. This does not add GGUF support, KV quantization, or active-context offloading.
Usage, raw benchmark results, and continuation limitations are documented in
docs/conversation-cache.md,docs/conversation-cache-results.md, anddocs/benchmarks/conversation-cache-m5.json.Disclosure: This implementation, tests, documentation, benchmarks, and PR description were produced with OpenAI Codex assistance. Codex also ran the reported local validation and prepared this draft PR at the repository owner's request.